Papers with deep understanding

4 papers
Don’t Judge a Book by its Cover: Testing LLMs’ Robustness Under Logical Obfuscation (2026.eacl-long)

Copied to clipboard

Challenge: obfuscated questions pose significant challenges for large language models . current models parse questions without deep understanding, MIT researchers say .
Approach: They propose a structure-preserving framework for logical obfuscation to test models . they use a logically equivalent framework to obliviate questions to logical equivalents .
Outcome: The proposed framework is a first-of-its-kind diagnostic benchmark with 1,108 questions . obfuscation severely degrades zero-shot performance, the authors show .
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks on longcontext large language models fail to reflect their deep understanding capabilities across diverse tasks.
Approach: They propose a benchmark to assess the ability of long-context large language models to handle long-text problems.
Outcome: The proposed model achieves 50.1% accuracy when directly answering the questions . human experts achieve only 53.7% accuracy under a 15-minute time constraint .
PaperRobot: Incremental Draft Generation of Scientific Ideas (P19-1)

Copied to clipboard

Challenge: a paper robot can read existing papers and create new nodes or links in the knowledge graphs.
Approach: They propose to automate the creation of new ideas by predicting links from the background KGs.
Outcome: The proposed paper automates three tasks: read existing papers, create new ideas, predict links . the paper generated abstracts, conclusion and future work sections, and new titles are chosen over human-written ones up to 30%, 24% and 12% of the time.
R-CHAR: A Metacognition-Driven Framework for Role-Playing in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing role-playing structures lack cognitive consistency in complex scenarios . Existing models excel in math and coding tasks but lack coherent reasoning .
Approach: They propose a metacognition-driven framework that enhances role-playing performance . experimental results show performance improvements across varying scenario complexities .
Outcome: The proposed framework outperforms existing models in social intelligence tasks and shows strength in long-context comprehension and group-level social interactions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations